Papers by Jong Inn Park
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these models reflect human-like cognition versus advanced pattern recognition remains an open question. |
| Approach: | They conduct a series of targeted experiments to assess whether LLMs construct semantic representations and pragmatic inferences in a human-like manner. |
| Outcome: | The proposed framework can be used to assess the cognitive and linguistic capabilities of large language models (LLMs). |
Benchmarking Cognitive Biases in Large Language Models as Evaluators (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been shown to be effective as automatic evaluators with simple prompting and in-context learning. |
| Approach: | They assemble 16 Large Language Models and evaluate their outputs by preference ranking . they introduce a cognitive bias benchmark to measure six different cognitive biases in LLM evaluation outputs. |
| Outcome: | The proposed model is biased on the CoBBLer benchmark, indicating that machine preferences are misaligned with humans. |
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs (2026.eacl-long)
Copied to clipboard
Karin de Langis, Jong Inn Park, Bin Hu, Khanh Chi Le, Andreas Schramm, Michael C. Mensink, Andrew Elfenbein, Dongyeop Kang
| Challenge: | Working memory is a critical component of human intelligence and executive functioning . it is correlated with performance on various cognitive tasks, including fluid intelligence . |
| Approach: | They apply working memory tasks to large language models to estimate working memory capacity . they find that LLMs exceed normative human scores, but not executive functioning benchmarks . |
| Outcome: | The proposed models do not show higher performance on executive functioning tasks or problem solving benchmarks. |